iT邦幫忙

2026 iThome 鐵人賽

DAY 6
0

現在會發生什麼

前五天我們建了一個完整的 Agent:讀檔 → 分塊 → 提取約束 → 檢測矛盾 → 生成報告。能用。

但試試用真實的 100 頁規格書,你會遇到四個問題:

問題 1:檢測規則太窄

Day 5 的 detect_by_rules() 只有 3 個硬編碼規則:多用戶 vs 單用戶、加密 vs 明文、性能數字。實際規格書中的矛盾模式有幾十種。新模式會被漏過。

問題 2:LLM 調用太慢

逐個 chunk 呼叫,逐對約束檢測。5 個 chunk 就要 5 次呼叫。200 個約束要 20000 次對比(O(n²))。如果用付費 API,成本會爆炸。

問題 3:沒有容錯

LLM 超時?整個分析失敗。格式錯誤?崩潰。

問題 4:沒有緩存

同一個 SRS 分析 10 次,做了 10 倍冤枉功。

三層修復

層 1:工具系統

不要硬編碼規則。定義可配置的工具,LLM 或用戶可以選擇調用哪些:

@dataclass
class Tool:
    name: str              # "extract_constraints"
    description: str       # 工具描述
    input_schema: dict     # JSON Schema (LLM 看這個決定怎麼調用)
    handler: Callable      # 實際執行函數

# 三個核心工具
TOOLS_REGISTRY = [
    Tool(name="extract_constraints", ...),
    Tool(name="detect_logical_conflict", ...),
    Tool(name="validate_constraint", ...),
]

優勢:新增工具不改代碼邏輯。只需定義新 Tool,加入註冊表。LLM 看到 JSON Schema,自動知道怎麼調用。

層 2:批處理與並行

不這樣:

for chunk in chunks:
    constraints += call_llm(f"提取 {chunk}")  # 5 次 API 呼叫

這樣:

with concurrent.futures.ThreadPoolExecutor(max_workers=3) as executor:
    results = executor.map(extract_chunk_constraints, chunks)
# 5 個 chunk 並行處理,實際只需 ~2 次 API 呼叫

結果:速度快 5-10 倍。

層 3:重試 + 緩存

超時會自動重試,退避等待:

@retry(max_attempts=3, backoff_factor=2.0)
def call_llm(prompt: str):
    ...

相同輸入緩存結果,避免重複計算:

@cache_result
def analyze_srs(srs_text: str):
    ...

工具系統實現

src/tools.py:

from dataclasses import dataclass
from typing import Callable, Dict, Any, List
from enum import Enum

class ToolType(str, Enum):
    EXTRACTION = "提取"
    DETECTION = "檢測"
    VALIDATION = "驗證"

@dataclass
class Tool:
    name: str
    description: str
    tool_type: ToolType
    input_schema: Dict[str, Any]
    handler: Callable

    def __call__(self, **kwargs) -> Any:
        return self.handler(**kwargs)

# 工具處理函數

def extract_constraints_handler(srs_text: str, chunk_id: int = 0) -> List[Dict]:
    """用正則表達式提取 REQ-xxx 需求"""
    pattern = r'(REQ-[\d.]+):([^(\n]+)'
    matches = re.findall(pattern, srs_text)
    return [
        {"id": req_id, "text": text.strip(), "chunk_id": chunk_id}
        for req_id, text in matches
    ]

def detect_logical_conflict_handler(constraint1: Dict, constraint2: Dict) -> Dict:
    """比對兩個約束,檢測邏輯矛盾"""
    text1 = constraint1.get("text", "")
    text2 = constraint2.get("text", "")

    # 規則 1:多用戶 vs 單用戶
    if ("多用戶" in text1 and "單用戶" in text2) or \
       ("多用戶" in text2 and "單用戶" in text1):
        return {
            "conflicting": True,
            "conflict_type": "邏輯矛盾",
            "severity": "高",
        }

    # ... 更多規則 ...
    return {"conflicting": False}

# 工具管理器
class ToolManager:
    def __init__(self, tools: List[Tool] = None):
        if tools is None:
            tools = TOOLS_REGISTRY
        self.tools = {tool.name: tool for tool in tools}

    def call_tool(self, tool_name: str, **kwargs) -> Any:
        tool = self.tools[tool_name]
        return tool(**kwargs)

    def list_tools(self) -> List[Dict]:
        """返回可用工具信息(給 LLM 看)"""
        return [
            {
                "name": tool.name,
                "description": tool.description,
                "input_schema": tool.input_schema,
            }
            for tool in self.tools.values()
        ]

tools_manager = ToolManager()

關鍵洞察:工具的 input_schema 是 JSON Schema。LLM 看著這個,就知道怎麼調用。你可以給 LLM 提供工具列表,LLM 自己決定什麼時候用哪個工具。這就是「工具使用」(tool use)。

優化:重試與緩存

重試裝飾器 — 失敗自動重試,等待時間指數增長(避免打爆服務):

def retry(max_attempts: int = 3, backoff_factor: float = 2.0):
    def decorator(func):
        @functools.wraps(func)
        def wrapper(*args, **kwargs):
            for attempt in range(1, max_attempts + 1):
                try:
                    return func(*args, **kwargs)
                except Exception as e:
                    if attempt < max_attempts:
                        wait_time = backoff_factor ** (attempt - 1)
                        print(f"⚠️  嘗試 {attempt} 失敗,等待 {wait_time:.1f} 秒...")
                        time.sleep(wait_time)
            raise last_exception
        return wrapper
    return decorator

# 使用
@retry(max_attempts=3, backoff_factor=1.5)
def call_ollama(prompt: str) -> str:
    response = requests.post(...)
    ...

超時時自動重試。第一次 0.5 秒等待,第二次 1.5 秒,第三次 3.5 秒。99% 的臨時故障會被吸收。

緩存裝飾器 — 相同輸入使用緩存結果,避免重複計算:

def cache_result(func):
    cache_dict = {}

    @functools.wraps(func)
    def wrapper(*args, **kwargs):
        # 生成緩存鍵(輸入的 MD5 哈希)
        key = hashlib.md5(
            json.dumps({"args": args, "kwargs": kwargs}, sort_keys=True)
        ).hexdigest()

        if key in cache_dict:
            print(f"💾 使用緩存({key[:8]}...)")
            return cache_dict[key]

        result = func(*args, **kwargs)
        cache_dict[key] = result
        return result

    return wrapper

# 使用
@cache_result
def analyze_srs(srs_text: str):
    ...
    # 同一個 SRS 分析多次,第二次立刻返回

並行化

用 concurrent.futures 並行處理多個 chunk:

def extract_chunk_constraints(chunk):
    return call_ollm(f"提取 {chunk.content}")

with concurrent.futures.ThreadPoolExecutor(max_workers=3) as executor:
    all_results = list(executor.map(extract_chunk_constraints, chunks))

constraints = [c for result in all_results for c in result]

5 個 chunk,3 個 worker,並行度 3。實際執行時間 ≈ 5/3 ≈ 2 次 LLM 呼叫的時間。

整合到 CLI

新增 --optimize 標誌:

@app.command()
def analyze_optimized(
    filepath: str = typer.Argument(...),
    cache: bool = typer.Option(True, "--cache/--no-cache"),
    retry: bool = typer.Option(True, "--retry/--no-retry"),):
    """使用優化版本分析 SRS"""
    try:
        agent = OllamaAgent()
        srs_text = agent.read_srs(filepath)

        if cache:
            from src.optimizations import analyze_srs_cached
            result = analyze_srs_cached(srs_text)
        else:
            from src.workflow import analyze_srs_with_workflow
            result = analyze_srs_with_workflow(srs_text)

        print(f"✅ 完成")
        print(f"  約束:{len(result['global_constraints'])}")
        print(f"  矛盾:{len(result['conflicts'])}")

        from src.tools import tools_manager
        print("\n可用工具:")
        for tool_info in tools_manager.list_tools():
            print(f"  • {tool_info['name']}: {tool_info['description']}")

    except Exception as e:
        print(f"❌ 錯誤: {e}")
        raise typer.Exit(1)

執行:

# 使用緩存和重試
python -m src.main analyze-optimized tests/fixtures/srs_medium.md

# 禁用緩存
python -m src.main analyze-optimized tests/fixtures/srs_medium.md --no-cache

為什麼這些改進有效

工具系統:擴展性。新增衝突模式?寫新工具,加入註冊表。不改現有代碼。

並行化:線性變指數。5 個任務從 5 次呼叫變 2 次。

重試:魯棒性。網絡抖動不再是致命傷。

緩存:效率。同一個 SRS 分析 10 次不會重複計算。

這些都是工程最佳實踐。個別看不起眼,組合起來性能會翻倍。

明天呢?

Day 7 我們會建評測框架:20+ 個標準 SRS 樣本,自動計算 precision/recall/F1,對比不同版本性能。

現在是「能用」的系統。明天開始,我們用數據驗證它有多好。


進度

第 1 週完成 ✅

  • Day 1-5:構建核心系統
  • Day 6:優化與容錯 ← 今天

第 2 週開始:評測與驗證

  • Day 7:評測框架
  • Day 8-10:性能測試與優化

上一篇
Day 5:把 State 變成能執行的流程
下一篇
Day 7:你怎麼知道它沒在騙你
系列文
解決需求規格書矛盾:用 Claude Code × MCP 實作自律型文檔審查 Agent 共 17 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言